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BACKGROUND OF THE INVENTION 
The present invention relates to the field of computer systems. More 
specifically, the present invention relates to computer systems for visualizing analysis 
results. 

Devices and computer systems for forming and using arrays of materials on 
a substrate are known. For example, PCT Publication No. WO 92/10588, incorporated 
herein by reference for all purposes, describes techniques for sequencing or sequence 
checking nucleic acids and other materials. Arrays for performing these operations may 
be formed according to the methods of, for example, the pioneering techniques disclosed 
in U.S. Patent No. 5,143,854 and U.S. Patent No. 5,593,839 both incorporated herem by 
reference for all purposes. 

According to one aspect of the techniques described therein, an array of 
nucleic acid probes is fabricated at known locations on a substrate or chip. A 
fluorescently labeled nucleic acid is then brought into contact with the chip and a scanner 
generates an image file (which is processed into a cell file) indicating the locations where 
the labeled nucleic acids bound to the chip. Based upon the cell file and identities of the 
probes at specific locations, it becomes possible to extract information such as the 
monomer sequence of DNA or RNA. Such systems have been used to form, for 
example, arrays of DNA that may be used to study and detect mutations relevant to cystic 
fibrosis, the P53 gene (relevant to certain cancers), HIV, and other genetic 
characteristics. 

Computer-aided techniques for monitoring gene expression using such 
arrays of probes have also been developed as disclosed in U.S. Patent Application No. 
08/828,952 (Attorney Docket No. 16528X-028900US) and PCT Publication No. 
WO 97/10365 (Attorney Docket No. 16528X-017110PC), the contents of which are 
herein incorporated by reference. Many disease states are characterized by differences in 
the expression levels of various genes either through changes in the copy number of the 
genetic DNA or through changes in levels of transcription {e.g, , through control of 
initiation, provision of RNA precursors, RNA processing, etc.) of particular genes. For 
example, losses and gains of genetic material play an important role in malignant 



transformation and progression. Furthermore, changes in the expression (transcription) 
levels of particular genes (e.g,, oncogenes or tumor suppressors), serve as signposts for 
the presence and progression of various cancers. 

It is desirable to identify genes having expression levels relevant to 
diagnosis of a diseased state by analyzing the expression levels of large numbers of genes 
in both diseased and normal individuals. Methods for collecting the expression level 
information have been developed. However, the user interfaces for gene expression 
monitoring systems that have been developed until now are designed to clearly present the 
expression of particular pre-selected genes. A user seeking to identify, e.g., an oncogene 
or a tumor suppressor gene, must individually review the expression level of large 
numbers of genes and compare the expression levels between diseased and normal 
individuals. What is needed is a user interface that takes advantage of collected gene 
expression information to help the user to identify particxilar genes of mterest. 

SUMMARY OF THE INVENTION 
The present invention provides innovative systems and methods for 
visualizing information collected from analyzing samples. The samples may include 
nucleic acids, proteins, or other polymers. Gene expression level as determined from 
analysis of a nucleic acid sample is one possible analysis result that may be visualized. In 
one embodiment, a computer system may display the expression levels of multiple genes 
simultaneously in a way that facilitates user identification of genes whose expression is 
significant to a characteristic such as disease or resistance to disease. Additionally, the 
computer system may facilitate display of further information about relevant genes once 
they are identified. 

A first aspect of the invention provides a computer-implemented method for 
presenting expression level information as collected from first and second samples. The 
method includes steps of: displaying a first axis corresponding to expression level m the 
first sample, and displaying a second axis substantially perpendicular to the first axis, the 
second axis corresponding to expression level in the second sample. The method further 
includes a step of: for a selected expressed sequence, displaying a mark at a position. 
The position is selected relative to the first axis in accordance with an expression level of 
the selected expressed sequence in the first sample and relative to the second axis in 
accordance with an expression level of the selected expressed sequence in the second 
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sample. A particularly useful application is displaying many marks simultaneously for 
many selected genes to discover which ones of the selected genes may be relevant to the 
characteristic. 

A second aspect of the invention provides a computer-implemented method 
5 of presenting sample analysis information. The method includes steps of: displaying a 
first axis corresponding to a concentration of a compound in a first sample as determined 
by monitoring binding of the compound to a selected polymer having binding affinity to 
the compound, and displaying a second axis substantially perpendicular to the first axis. 
The second axis corresponds to a concentration of the compound in the second sample as 
10 determined by monitoring binding of the compound to the selected polymer. The method 

further preferably includes a step of displaying a mark at a position. The position is 
O selected relative to the first axis in accordance with the concentration in the first sample 
3 and relative to the second axis in accordance with the concentration in the second sample. 

A further understanding of the nature and advantages of the inventions 
herein may be realized by reference to the remaining portions of the specification and the 
attached drawings. 

m BRIEF DESCRIPTION OF THE DRAWINGS 

3 Fig. 1 illustrates an example of a computer system that may be used to 

Jo execute software embodiments of the present invention. 

Fig. 2 shows a system block diagram of a typical computer system. 

Fig. 3 illustrates an overall system for forming and analyzing arrays of 
polymers including biological materials such as DNA or RNA. 

Fig. 4 is an illustration of an embodiment of software for the overall 

25 system. 

Fig. 5 shows a flowchart of a process of monitoring the expression of a 
gene by comparing hybridization intensities of pairs of perfect match and mismatch 
probes. 

Fig. 6 shows a screen display illustrating gene expression levels for 
30 multiple genes as collected from both normal and diseased tissue. 

Figs. 7A-7B show screen displays illustrating information about a particular 
gene selected from the display of Fig. 6. 
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DESCRIPTION OF SPECIFIC EMBODIMENTS 
The present invention provides innovative methods of monitoring 
visualizing gene expression. In the description that follows, the invention will be 
described in reference to preferred embodiments. However, the description is provided 
5 for purposes of illustration and not for limiting the spkit and scope of the invention. 

Fig. 1 illustrates an example of a computer system that may be used to 
execute software embodiments of the present invention. Fig. 1 shows a computer system 
1 which includes a monitor 3, screen 5, cabinet 7, keyboard 9, and mouse 11. Mouse 11 
may have one or more buttons such as mouse buttons 13, Cabinet 7 houses a CD-ROM 
10 drive 15 and a hard drive (not shown) that may be utilized to store and retrieve software 
programs including computer code incorporating the present invention. Although a CD- 
Q ROM 17 is shown as the computer readable medium, other computer readable media 
n inchxdmg floppy disks, DRAM, hard drives, flash memory, tape, and the like may be 
2:! utilized. Cabinet 7 also houses familiar conqniter components (not shown) such as a 
'15 processor, memory, and the like. 

Fig. 2 shows a system block diagram of computer system 1 used to execute 
L. software embodiments of the present invention. As m Fig, 1, computer system 1 includes 

monitor 3 and keyboard 9. Computer system 1 fiirther includes subsystems such as a 
3 central processor 50, system memory 52, I/O controller 54, display adapter 56, 
;3o removable disk 58, fixed disk 60, network interface 62, and speaker 64. Removable disk 
58 is representative of removable computer readable media like floppies, tape, CD-ROM, 
removable hard drive, flash memory, and the like. Fixed disk 60 is representative of an 
internal hard drive or the like. Other computer systems suitable for use with the present 
invention may include additional or fewer subsystems. For example, another computer 
25 system could include more than one processor 50 {Le, , a multi-processor system) or 
memory cache. 

Arrows such as 66 represent the system bus architecture of computer 
system 1. However, these arrows are illustrative of any interconnection scheme serving 
to link the subsystems. For example, display adapter 56 may be connected to central 
30 processor 50 through a local bus or the system may include a memory cache. Computer 
system 1 shown in Fig. 2 is but an example of a computer system suitable for use with 
the present invention. Other configurations of subsystems suitable for use with the 
present invention will be readily apparent to one of ordinary skill in the art. In one 
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embodiment, the computer system is an IBM compatible personal computer. 

The VLSIPS"^ and GeneChip™ technologies provide methods of making and 
using very large arrays of polymers, such as nucleic acids, on very small chips. See U.S. 
Patent No. 5,143,854 and PCX Patent Publication Nos. WO 90/15070 and 92/10092, 
5 each of which is hereby incorporated by reference for all purposes. Nucleic acid probes 
on the chip are used to detect complementary nucleic acid sequences in a sample nucleic 
acid of interest (the "target" nucleic acid). 

It should be understood that the probes need not be nucleic acid probes but 
may also be other receptors, such as antibodies, or polymers such as peptides. Peptide 
10 probes may be used to detect the concentration of other peptides, proteins, or other 

compounds in a sample. The probes must be carefiilly selected to have bonding affinity 
O to the compound whose concentration they are to be used to measure. 
2 In one embodiment, the present invention provides methods of visualizing 

2f mformation relating to the concentration of compounds in a sample as measured by 
^ monitoring affinity of the compounds to probes. In a particular application, the 
yj concentration information is generated by analysis of hybridization intensity files for a 

chip containing hybridized nucleic acid probes. The hybridization of a nucleic acid 
fll sample to certain probes may represent the expression level of one more genes or 
51 expressed sequence tags (ESTs). The expression level of a gene or EST is herein 
'Jo imderstood to be the concentration within a sample of mRNA or protein that would result 
from the transcription of the gene or EST. 

Expression level information visualized by virtue of the present invention 
need not be obtained from probes but may originate from any source. If the expression 
information is collected from a probe array, the probe array need not meet any particular 
25 criteria for size and density. Furthermore, the present invention is not limited to 
visualizing fluorescent measurements of bondings such as hybridizations but may be 
readily utilized to visualize other measurements. 

Concentration of compounds other than nucleic acids may be visualized 
according to one embodiment of the present invention. For example, a probe array may 
30 include peptide probes which may be exposed to protein samples, polypeptide samples, or 
other compounds which may or may not bond to the peptide probes. By appropriate 
selection of the peptide probes, one may detect the presence or absence of particular 
compounds which would bond to the peptide probes. 
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For purposes of illustration, the present invention is described as being part 
of a system that designs a chip mask, synthesizes the probes on the chip, labels nucleic 
acids from a target sample, and scans the hybridized probes. Such a system is set forth 
in U.S. Patent No. 5,571,639 which is hereby incorporated by reference for all purposes. 
5 However, the present invention may be used separately from the overall system for 

analyzing data generated by such systems, such as at remote locations, or for visualizing 
the results of other systems for generating expression information, or for visualizing 
concentrations of polymers other than nucleic acids. 

Fig. 3 illustrates a computerized system for forming and analyzing arrays 
10 of biological materials such as RNA or DNA. A computer 100 is used to design arrays 
of biological polymers such as RNA or DNA. The computer 100 may be, for example, 
O an appropriately programmed IBM personal computer compatible running Windows NT 
^ including appropriate memory and a CPU as shown in Figs. 1 and 2. The computer 

system 100 obtains inputs from a user regarding characteristics of a gene of interest, and 
1^ other inputs regarding the desired features of the array. Optionally, the computer system 
iJi may obtain information regarding a specific genetic sequence of interest from an extemal 
1,^ or internal database 102 such as GenBank. The output of the computer system 100 is a 
HJ set of chip design computer files 104 in the form of, for example, a switch matrix, as 
S described in PCT application WO 92/10092, and other associated computer files. 
^ The chip design files are provided to a system 106 that designs the 

lithographic masks used in the fabrication of arrays of molecules such as DNA. The 
system or process 106 may include the hardware necessary to manufacture masks 110 and 
also the necessary computer hardware and software 108 necessary to lay the mask 
patterns out on the mask in an efficient manner. As with the other features in Fig. 3, 
25 such equipment may or may not be located at the same physical site, but is shown 

together for ease of illustration in Fig. 3. The system 106 generates masks 110 or other 
synthesis patterns such as chrome-on-glass masks for use in the fabrication of polymer 
arrays. 

The masks 110, as well as selected information relating to the design of the 
30 chips from system 100, are used in a synthesis system 112. Synthesis system 112 

includes the necessary hardware and software used to fabricate arrays of polymers on a 
substrate or chip 114. For example, synthesizer 112 includes a light source 116 and a 
chemical flow cell 118 on which the substrate or chip 114 is placed. Mask 110 is placed 
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between the light source and the substrate/chip, and the two are translated relative to each 
other at appropriate times for deprotection of selected regions of the chip. Selected 
chemical reagents are directed through flow cell 118 for coupling to deprotected regions, 
as well as for washing and other operations. All operations are preferably directed by an 
5 appropriately programmed computer 119, which may or may not be the same computer as 
the computer (s) used in mask design and mask making. 

The substrates fabricated by synthesis system 112 are optionally diced into 
smaller chips and exposed to marked targets. The targets may or may not be 
complementary to one or more of the molecules on the substrate. The targets are marked 
10 with a label such as a fluorescein label (indicated by an asterisk in Fig. 3) and placed in 

scanning system 120. Scanning system 120 again operates under the direction of an 
□ appropriately programmed digital computer 122, which also may or may not be the same 
S computer as the computers used in synthesis, mask making, and mask design. The 
^ scanner 120 includes a detection device 124 such as a confocal microscope or CCD 
'15 (charge-coupled device) that is used to detect the location where labeled target has bound 
to the substrate. The output of scanner 120 is an image flle(s) 124 indicating, in the case 
t,^ of fluorescein labeled target, the fluorescence intensity (photon counts or other related 
fU measurements, such as voltage) as a function of position on the substrate. Since higher 
"5 photon counts will be observed where the labeled target has bound more strongly to the 
^ array of polymers, and since the monomer sequence of the polymers on the substrate is 
known as a function of position, it becomes possible to determine the sequence(s) of 
polymer(s) on the substrate that are complementary to the target. 

The image file 124 is provided as input to an analysis system 126 that 
incorporates the visualization and analysis methods of the present invention. Again, the 
25 analysis system may be any one of a wide variety of computer system. The present 

invention provides various methods of analyzing and visualizing the chip design files and 
the image files, providing appropriate output 128. The chip design need not include any 
particular number of probes. It should be understood that the present invention does not 
require any particular source of expression level information. 
30 Fig. 4 provides a simplified illustration of the overall software system used 

in the operation of one embodiment of the invention. As shown in Fig. 4, the system 
first identifies the nucleotide sequence(s) or targets that would be of interest in a 
particular expression level analysis at step 202. The sequences of interest correspond to 



8 

mRNA transcripts of one or more genes, ESTs or nucleic acids derived from the mRNA 
transcripts. Sequence selection may be provided via manual input of text files or may be 
from external sources such as GenBank. 

At step 204 the system evaluates the sequences of interest to determine or 
5 assist the user in determining which probes would be desirable on the chip, and provides 
an appropriate "layout" on the chip for the probes. The process of selecting probes for 
an expression level analysis is explained in PCT Publication No. WO 97/10365, the 
contents of which are herein incorporated by reference. An alternative probe selection 
process that does not require prior knowledge of sequences of interest is explained in 
10 PCT Publication No. W097/27317 (Attomey Docket No. 18547-01 94 lOPC), the contents 

of which are herein incorporated by reference. Further general background on probe 
□ selection is found in PCT Publication No. W095/11995 (Attomey Docket No. 18547- 
g 00411 IPC) and PCT Publication No. W097/29212 (Attomey Docket No. 18547- 
^ 018540PC), the contents of which are herein incorporated by reference. The term 
1^ "perfect match probe" refers to a probe that has a sequence that is perfectly 
cl complementary to a particular target sequence. The test probe is typically perfectly 

complementary to a portion (subsequence) of the target sequence. The term "mismatch 

Q 

IFU control" or "mismatch probe" refer to probes whose sequence is deliberately selected not 
!^ to be perfectly complementary to a particular target sequence. For each mismatch (MM) 
^ control in an array there typically exists a corresponding perfect match (PM) probe that is 
perfectly complementary to the same particular target sequence. 

The process compares hybridization intensities of pairs of perfect match 
and mismatch probes that are preferably covalently attached to the surface of a substrate 
or chip. Most preferably, the nucleic acid probes have a density greater than about 60 
25 different nucleic acid probes per 1 cm-^ of the substrate. 

Initially, nucleic acid probes are selected that are complementary to the 
target sequence. These probes are the perfect match probes. Another set of probes is 
specified that are intended to be not perfectly complementary to the target sequence. 
These probes are the mismatch probes and each mismatch probe includes at least one 
30 nucleotide mismatch from a perfect match probe. Accordingly, a mismatch probe and the 
perfect match probe to which it is identical except for one base make up a pair. As 
mentioned earlier, the nucleotide mismatch is preferably near the center of the mismatch 
probe. 
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The probe lengths of the perfect match probes are typically chosen to 
exhibit detectably greater hybridization with the target sequence relative to the mismatch 
probes. For example, the nucleic acid probes may be all 20-mers. However, probes of 
varying lengths may also be synthesized on the substrate for any number of reasons 
5 including resolving ambiguities. 

Again referring to Fig. 4, at step 206 the masks for the synthesis are 
designed. At step 208 the software utilizes the mask design and layout information to 
make the DNA or other polymer chips. This step 208 will control, among other things, 
relative translation of a substrate and the mask, the flow of desired reagents through a 

10 flow cell, the synthesis temperature of the flow cell, and other parameters. At step 210, 
another piece of software is used in scanning a chip thus synthesized and exposed to a 
labeled target. The software controls the scanning of the chip, and stores the data thus 

^ obtained in a file that may later be utilized to extract hybridization information. 

W At step 212 a computer system utilizes the layout information and the 

€§ fluorescence information to evaluate the hybridized nucleic acid probes on the chip. 
Among the important pieces of information obtained from DNA chips are the relative 

1^ fluorescent intensities obtained from the perfect match probes and mismatch probes. 

fu These intensity levels are used to estimate an expression level for a gene or EST. The 
computer system used for analysis will preferably have available other details of the 

#) experiment including possibly the gene name, gene sequence, probe sequences, probe 
locations on the substrate, and the like. 

According to the present invention, at step 214, the same computer system 
used for analysis or another one displays the expression level information in a format 
useful for identifying genes of interest. The visualized expression level information may 

25 include information collected from multiple applications of one or more previous steps of 
Fig. 4. 

Fig. 5 is a flowchart describing steps of estimating an expression level for 
a particular gene and determining whether the expression level is sufficiently high to be 
displayed. At step 952, the computer system receives raw scan data of N pairs of perfect 
30 match and mismatch probes. In a preferred embodiment, the hybridization iatensities are 
photon counts from a fluorescein labeled target that has hybridized to the probes on the 
substrate. For simplicity, the hybridization intensity of a perfect match probe will be 
designed 'T " and the hybridization intensity of a mismatch probe will be designed 
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*inm- 

Hybridization intensities for a pair of probes are retrieved at step 954, The 
background signal intensity is subtracted from each of the hybridization intensities of the 
pair at step 956. Background subtraction can also be performed on all the raw scan data 
5 at the same time. 

At step 958, the hybridization intensities of the pair of probes are compared 
to a difference threshold (D) and a ratio threshold (R), It is determined if the difference 
between the hybridization intensities of the pair (Ip^ - Ij^) is greater than or equal to the 
; , " difference threshold AND the quotient of the hybridization intensities of the pair (I^^ I 
10 Ij^jj^) is greater than or equal to the ratio threshold. The difference thresholds are 

typically user defined values that have been determined to produce accurate expression 
O monitoring of a gene or genes. In one embodiment, the difference threshold is 20 and the 
'p. ratio threshold is 1.2. 

If Ipjj^ - > = D and Ip^^ / > = R, the value NPOS is mcremented 
15 at step 960, In general, NPOS is a value that indicates the nimiber of pairs of probes 

which have hybridization intensities indicating that the gene is likely expressed, NPOS is 
1^^^ utilized in a determination of the expression of the gene. 

m At step 962, it is determined if lrma^\m>^^ and Ij^^ / Ipm > = If 

'5 these expressions are true, the value NNEG is incremented at step 964. In general, 
'"Mi NNEG is a value that indicates the number of pairs of probes which have hybridization 
intensities indicating that the gene is likely not expressed. NNEG, like NPOS, is utilized 
in a determination of the expression of the gene. 

For each pair that exhibits hybridization intensities either indicating the 
gene is expressed or not expressed, a log ratio value (LR) and intensity difference value 
25 (IDIF) are calculated at step 966. LR is calculated by the log of the quotient of the 

hybridization intensities of the pair (Ipj^ / 1^^^). The IDIF is calculated by the difference 
between the hybridization intensities of the pair (l^^ - 1^^)- If there is a next pair of 
hybridization intensities at step 968, they are retrieved at step 954. 

At step 972, a decision matrix is utilized to indicate if the gene is 
30 expressed. The decision matrix utilizes the values N, NPOS, NNEG, LR (multiple LRs), 
and IDIF (multiple IDIFs). The following four assignments are performed: 
PI = NPOS / NNEG 
P2 = NPOS / N 
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P3 = SUM(LR) / N 

P4 = SUM(IDIF)/N 
These P values are then utilized to determine if the gene is expressed and if the 
expression level should be displayed. In a preferred embodiment, the expression level of 
5 a gene should be displayed if: 

PI > 2.2 

P2 > 0.3 

P3 > 0.8 

P4 > 30 

10 Once all the pairs of probes have been processed and the expression of the 

gene indicated, an average of the IDIF values for the probes that incremented NPOS or 
p NNEG is calculated at step 975, which is utilized as an expression level. Of course, 
S other values including one of PI through P4 could be used to indicate expression level. 
^ For simplicity. Fig. 5 was described in reference to a single gene or EST. 

15 However, the visualization system of the present invention displays expression results for 
77- many genes to facilitate discovery of genes of interest or ESTs. Furthermore, the present 
1 invention contemplates display of expression levels of a single gene or ESTs as collected 
7- from two or more different samples such as tissue samples. The sample sources 
preferably differ in some characteristic. It will be understood that when the term 
#0 "sample" is used herein, measurements made on a single "sample" can be based on an 
aggregation of multiple sample collection events or even multiple organisms. 

Fig. 6 shows a screen display illustrating gene expression levels for 
multiple genes as collected from two tissue samples. A displayed horizontal axis 1002 
represents expression level measured in one or more nucleic acid samples taken from the 
25 first tissue sample. A displayed vertical axis 1004 represents expression level in one or 
more nucleic acid samples taken from the second tissue sample. Each of marks 1006 
represent a particular gene whose expression level has been measured in both the first and 
second tissue samples. Each mark 1006 is placed at a distance from vertical axis 1004 
corresponding to expression level in the first tissue sample and at a distance from the 
30 horizontal axis 1002 corresponding to expression level in the second tissue sample. 

The expression levels used for determining the position of marks 1006 are 
preferably taken from the result of step 975. The position of each of marks 1006 depends 
on two iterations of the steps of Fig. 5, once for the sample taken from the first tissue 
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sample and once for the sample taken from the second tissue sample. However, a mark 
is preferably displayed only if one of the samples meets the threshold criteria at step 972. 

In the depicted representative screen display, the first tissue sample is a 
cancerous tissue sample and the second tissue sample is a normal tissue sample. The 
5 individual marks represent the expression levels of selected genes in both cancerous and 
normal tissue. A first group of marks 1008 represent genes that are neither tumor 
suppressors nor oncogenes since their expression levels are roughly similar for both 
normal and cancerous tissue. These marks 1008 fall roughly along a line which is rotated 
45 degrees from each of the axes. A second group of marks 1010 represent genes that 

10 are likely oncogenes since their expression levels are found to be significantly higher in 
cancerous tissue than in normal tissue. A third group of marks 1012 represent genes that 

g are likely tumor suppressors since their expression levels are found to be significantly 
higher in normal tissue than in cancerous tissue. It will be appreciated that expression 

?u levels for large numbers of genes can be reviewed at once to discover the oncogenes and 

O txmior suppressors. 

rj: Although in the depicted display, the two types of tissue are normal tissue 

f and cancerous tissue, the present invention would aid in the discovery of genes whose 
fy expression is associated with any characteristic that varies among tissue samples. For 
example, once can compare expression results from tissue from individuals who have 
#) been exposed to HIV but remain infected to tissue obtained from infected individuals to 
identify genes conferring resistance to HIV. One can compare expression results between 
tissue from plants that survive drought to plants that do not. One can compare expression 
levels among tissue samples at successive stages or severity levels of the same disease, 
among tissue samples where different ultimate outcomes of the disease (e.g., patient death 
25 or remission) are known, among diseased tissue samples that have been subject to 
different treatment regimes including e.g, chemotherapy, antisense RNA, etc. For 
cancers, one can compare expression levels between malignant cells and non-malignant 
cells. Also expression levels can be compared among different organs, between species, 
and among different stages of development of an organ. 
30 It will be appreciated that the present invention also encompasses displays 

with more than two dimensions. A third visual dimension can be used to illustrate 
expression level from a third tissue sample. The time dimension can also be used to 
illustrate successive groups of two or three tissue samples at successive time periods. 
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The time dimension can be also used to correspond to tissue samples obtained at, e.g, 

successive stages of a disease. 

Other interface methods corresponding to human senses other than sight 

can also be incorporated within the presentation system of the present invention. The 
5 senses may correspond to additional dimensions. For example, marks can be displayed in 

succession accompanies by a sound having characteristics corresponding to expression 

level in another tissue sample. 

The user can employ a cursor 1014 to identify a particular mark as being 

of interest. Cursor 1014 can be moved to a particular mark by use of, e.g., mouse 11, 
10 Once cursor 1014 is over a mark of interest, the mark can be selected by, e.g., 

depression of one of mouse buttons 13. Selection of a particular mark can be facilitated 
Q by use of a zoom display feature (not shown). Once a particular mark is selected, further 

information is displayed about the gene represented by the mark. A special mouse can 

transmit a tactile sensation back to the user corresponding to expression level in a tissue 
45 sample as the user passes the mouse over a corresponding mark. 

2 It will be appreciated that the display of Fig. 6 is not limited to expression 

information. The two dimensions of Fig. 6 may correspond to indicators of the presence 
ffj of various polymers other than nucleic acids in two different samples. For example, each 
'% mark may correspond to a different polymer, polypeptide, or other compound. The 
#0 distance of the mark from each axis would correspond to a measure of presence of the 
particular polymer in the sample corresponding to the axis. One possible measure is 
produced by fluorescently tagging polymer samples such as protein samples and exposing 
a probe array such as a peptide probe array to the protein samples. The fluorescent 
intensity of the probes will then correspond to the bonding affinity of the sample to the 
25 probes. The intensity measurement or a measurement derived from the intensity 
measurement may then be used to position the marks of Fig. 6. 

Fig. 7 A shows a screen display giving information about a particular gene 
selected from the display of Fig. 6. A cluster number 702, a GenBank accession number 
704, and a verbal description 706 for the selected gene are displayed. The user can also 
30 select a number of marks 1006 by circling them with cursor 1014. Then a list of 
information as shown in Fig. 7 A is displayed for all the genes corresponding to the 
selected marks. 

By selecting GenBank accession nmnber 704 with another cursor (not 
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shown), the user can direct retrieval of the GenBank information for the selected gene. If 
the GenBank information is not available locally, the retrieval process can include 
formulating a query and transmitting the query to a GenBank web site. Once the 
GenBank information is retrieved, it can also be displayed. Fig. 7B depicts the GenBank 
5 information for the gene identified in Fig. 7A. 

In the foregoing specification, the invention has been described with 
reference to specific exemplary embodiments thereof. It will, however, be evident that 
various modifications and changes may be made thereunto without departing from the 
broader spirit and scope of the invention as set forth in the appended claims and their full 
10 scope of equivalents. 
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WHAT IS CLAIMED IS: 



1 1. A computer-implemented method of presenting expression level 

2 information as collected from first and second samples, said method comprising the steps 

3 of: 

4 displaying a first axis corresponding to expression level in said first 

5 sample; 

6 displaying a second axis substantially perpendicular to said first axis, said 

7 second axis corresponding to expression level in said second sample; and 

8 for a selected expressed sequence, displaying a mark at a position, wherein 

9 said position is selected relative to said first axis in accordance with an expression level 
10 of said selected expressed sequence in said first sample and relative to said second axis in 
ffl accordance with an expression level of said selected expressed sequence in said second 
ft sample. 



SI 2. The method of claim 1 wherein said selected expressed sequence 

JI2 comprises a gene. 

fl^l 3. The method of claim 1 wherein said selected expressed sequence 

jS2 comprises a portion of a gene. 

1 4. The method of claun 1 further comprising title step of repeating said 

2 displaying a mark step for a plurality of selected expressed sequences. 

1 5. The method of claim 1 further comprising the steps of: 

2 monitoring said expression level of said expressed sequence in said first 

3 sample and said second sample. 

1 6. The method of claim 3 wherein said monitormg step for one of said 

2 samples comprises substeps of: 

3 inputting a plurality of hybridization intensities of pairs of perfect match 

4 and mismatch probes, said perfect match probes being perfectly complementary to a 

5 target nucleic acid sequence indicative of expression of said selected gene and said 

6 mismatch probes having at least one base mismatch with said target sequence, and said 
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7 hybridization intensities indicating hybridization affinity between said perfect match and 

8 mismatch probes and a sample nucleic acid sequence from said one of said samples; 

9 comparing the hybridization intensities of each pair of perfect match probe 

10 and mismatch probe; and 

1 1 generating said expression level for said expressed sequence and said one 

12 of said samples responsive to results of said comparing step. 

1 7. The method of claim 6 further comprising the step of: 

2 comparing a difference between hybridization intensities of perfect match 

3 and mismatch probes at a base position to a difference threshold. 

CH 8. The method of claim 7 further comprising the step of: 

comparing a quotient of hybridization intensities of perfect match and 
mismatch probes at a base position to a ratio threshold. 

%. 2 

9. The method of claim 6 further comprising the steps of: 
iJ, a) counting a probe pair as a positive probe pair to increment a 

fl3 positive probe pair count if a perfect match probe intensity minus a mismatch probe 
Jfk intensity exceeds a difference threshold and said perfect match probe intensity divided by 
said mismatch probe intensity exceeds a ratio threshold; 

6 b) counting said probe pair as a negative probe pair to increment a 

7 negative probe pair count if said mismatch probe intensity minus said perfect match probe 

8 intensity exceeds said difference threshold and said mismatch probe intensity divided by 

9 said perfect match probe intensity exceeds said ratio threshold; and 

10 c) computing a logarithmic ratio of said perfect match probe intensity 

11 to said mismatch probe intensity. 

1 10. The method of claim 9 further comprising the steps of: 

2 repeating said a), b), and c) steps for each of said probe pairs, 

3 accximulating a sum of differences of said perfect match and mismatch probe intensities 

4 for probe pairs that cause; and 

5 determining an expression level of said selected expressed sequence to be 

6 an average of said differences. 
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1 11, The method of claim 1 further comprising the steps of: 

2 receiving user input selecting said mark; and 

3 in response to said user input, displaying information about said selected 

4 expressed sequence. 

1 12. The method of claim 11 further comprising the steps of: 

2 in response to said user input, displaying information about said selected 

3 expressed sequence. 

1 13. The method of claim 12 wherein said information about said 

2 selected expressed sequence comprises a GenBank accession number. 

% 14. The method of claim 12 wherein said information about said 

J selected expressed sequence comprises a GenBank database record for said selected 

's$ expressed sequence. 

^,1 15. The method of claim 1 wherein said first sample and said second 

raE sample are collected from tissue samples differing in a particular characteristic. 

£^ 16. The method of claim 15 wherein said particular characteristic 

2 comprises presence of disease. 

1 17. The method of claim 15 wherein said particular characteristic 

2 comprises a treatment strategy for a disease. 

1 18. The method of claim 1 wherein said particular characteristic is a 

2 stage of a disease. 

1 19. The method of claim 1 further comprising the step of : 

2 displaying a third axis substantially perpendicular to said first axis and to 

3 said second axis in a three-dimensional display environment wherein said position of said 

4 mark is further selected relative to said third axis in accordance with an expression level 

5 of said selected expressed sequence in a third sample. 
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1 20. A computer-implemented method of presenting sample analysis 

2 information comprising the steps of: 

3 displaying a first axis corresponding to a concentration of a compound m a 

4 first sample as determined by monitoring binding of said compound to a selected polymer 

5 having binding affinity to said compound; 

6 displaying a second axis substantially perpendicular to said first axis, said 

7 second axis corresponding to a concentration of said compound in said second sample as 

8 determined by monitoring binding of said compound to said selected polymer; and 

9 displaying a mark at a position, wherein said position is selected relative to 

10 said first axis in accordance with said concentration in said first sample and relative to 

11 said second axis in accordance with said concentration in said second sample. 

% 21. The method of claim 20 wherein said selected polymer comprises a 

^ nucleic acid sequence. 

Jl 22. The method of claim 20 wherein said selected polymer comprises a 

zj^ protein. 

23. The method of claim 21 further comprising the step of: 
% obtaining said concentration of said compound in said first sample by 

3 exposing said first sample to a plurality of nucleic acid probes. 

1 24. The method of claim 22 further comprising the step of: 

2 obtaining said concentration of said compound in said first sample by 

3 exposing said first sample to a plurality of peptide probes. 

1 25. A computer program product for presenting expression level 

2 information as collected from first and second samples, said product comprising:: 

3 code for displaying a first axis corresponding to expression level in said 

4 first sample; 

5 code for displaying a second axis substantially perpendicular to said first 

6 axis, said second axis corresponding to expression level in said second sample; 

7 code for, for a selected expressed sequence, displaying a mark at a 
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8 position, wherein said position is selected relative to said first axis in accordance with an 

9 expression level of said selected expressed sequence in said first sample and relative to 

10 said second axis in accordance with an expression level of said selected expressed 

11 sequence in said second sample; and 

12 a computer-readable storage medium for storing the codes. 

1 26. The product of claim 25 wherein said selected expressed sequence 

2 comprises a gene. 

1 27. The product of claim 25 wherein said selected expressed sequence 

2 comprises a portion of a gene. 

2 28. The product of claim 25 further comprising code for repeatedly 

S applying said displaying a mark code for a plurality of selected expressed sequences. 

29. The product of claim 25 further comprising: 
12 code for monitoring said expression level of said expressed sequence in 

ni said first sample and said second sample. 

. ir'h 

j|l 30. The product of claim 27 wherein said monitoring step for one of 

2 said samples comprises: 

3 code for inputting a plurality of hybridization intensities of pairs of perfect 

4 match and mismatch probes, said perfect match probes being perfectly complementary to 

5 a target nucleic acid sequence indicative of expression of said selected gene and said 

6 mismatch probes having at least one base mismatch with said target sequence, and said 

7 hybridization intensities indicatmg hybridization affinity between said perfect match and 

8 mismatch probes and a sample nucleic acid sequence from said one of said samples; 



9 comparing the hybridization intensities of each pair of perfect match probe 

10 and mismatch probe; and 

11 generating said expression level for said expressed sequence and said one 

12 of said samples responsive to results of said comparing step. 
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1 31. The product of claim 30 further comprising: 

2 code for comparing a difference between hybridization intensities of perfect 

3 match and mismatch probes at a base position to a difference threshold. 

1 32. The product of claim 31 further comprising: 

2 code for comparing a quotient of hybridization intensities of perfect match 

3 £ind mismatch probes at a base position to a ratio threshold. 

1 33. The product of claim 30 further comprising: 

2 a) code for counting a probe pair as a positive probe pair to increment a 

3 positive probe pah count if a perfect match probe intensity minus a mismatch probe 

il intensity exceeds a difference threshold and said perfect match probe intensity divided by 

M said mismatch probe intensity exceeds a ratio threshold; 

l| b) code for counting said probe pair as a negative probe pair to increment 

'9 a negative probe pair count if said mismatch probe intensity minus said perfect match 

probe intensity exceeds said difference threshold and said mismatch probe intensity 

'z^ divided by said perfect match probe intensity exceeds said ratio threshold; and 

c) code for computing a logarithmic ratio of said perfect match probe 

fcl intensity to said mismatch probe intensity, 

1 34. The product of claim 33 further comprising: 

2 code for repeatedly applyhig said a), b), and c) codes for each of said 

3 probe pairs, accumulating a sum of differences of said perfect match and mismatch probe 

4 intensities for probe pairs that cause; and 

5 code for determining an expression level of said selected expressed 

6 sequence to be an average of said differences. 

1 35. The product of claim 25 further comprising: 

2 code for receiving user input selecting said mark; and 

3 code for, in response to said user input, displaying information about said 

4 selected expressed sequence. 
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1 36. The product of claim 35 ftirther comprising: 

2 code for, in response to said user input, displaying information about said 

3 selected expressed sequence. 

1 37. The product of claim 36 wherein said information about said 

2 selected expressed sequence comprises a GenBank accession number. 

1 38. The product of claim 36 wherein said information about said 

2 selected expressed sequence comprises a GenBank database record for said selected 

3 expressed sequence. 

^ 39. The product of claim 25 wherein said first sample and said second 

S sample are collected from tissue samples differing in a particular characteristic. 

-If 40. The product of claim 39 wherehi said particular characteristic 

|2 comprises presence of disease. 

ni 41. The product of claim 39 wherein said particular characteristic 

% comprises a treatment strategy for a disease. 

1 42. The product of claim 25 wherein said particular characteristic is a 

2 stage of a disease. 

1 43. The product of claim 25 further comprising the step of : 

2 displaying a third axis substantially perpendicular to said first axis and to 

3 said second axis in a three-dimensional display environment wherein said position of said 

4 mark is further selected relative to said third axis in accordance with an expression level 

5 of said selected expressed sequence in a third sample. 

1 44. A computer program product for presenting sample analysis 

2 information comprising: 

3 code for displaying a first axis corresponding to a concentration of a 

4 compound in a first sample as determined by monitoring binding of said compovmd to a 
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5 selected polymer having bonding affinity to said compoxmd; 

6 code for displaying a second axis substantially perpendicular to said first 

7 axis, said second axis corresponding to concentration of said compound in a second 

8 sample as determined by monitoring binding of said compound to said selected polymer; 

9 code for displaying a mark at a position, wherein said position is selected 

10 relative to said first axis in accordance with said concentration in said first sample and 

11 relative to said second axis in accordance with said concentration in said second sample; 

12 and 

13 a computer-readable storage medium that stores the codes. 

1 45. The product of claim 44 wherein said selected polymer comprises a 

g| nucleic acid sequence. 

J 46. The product of claim 44 wherein said selected polymer comprises a 

i protein, 

47. A computer system comprising a display, a processor, and a 

H memory that stores histructions for configuring said processor to: 
% display a first axis corresponding to expression level in said first sample; 

display a second axis substantially perpendicular to said first axis, said 

5 second axis corresponding to expression level in said second sample; and 

6 for a selected expressed sequence, display a mark at a position, wherein 



7 said position is selected relative to said first axis in accordance with an expression level 

8 of said selected expressed sequence in said first sample and relative to said second axis in 

9 accordance with an expression level of said selected expressed sequence in said second 
10 sample. 



1 48. A computer system comprising a display, a processor, and a 

2 memory that stores instructions for configuring said processor to: 

3 display a first axis corresponding to a concentration of a compound in a 

4 first sample as determined by monitoring buiding of said compoxmd to a selected polymer 

5 having binding affinity to said compound; 

6 display a second axis substantially perpendicular to said first axis, said 
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7 second axis corresponding to a concentration of said compound in said second sample as 

8 determined by monitoring binding of said compound to said selected polymer; and 

9 display a mark at a position, wherein said position is selected relative to 

10 said first axis in accordance with said concentration in said first sample and relative to 

11 said second axis in accordance with said concentration in said second sample. 
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COMPUTER-AIDED VISUALIZATION 
OF EXPRESSION COMPARISON 



ABSTRACT OF THE DISCLOSURE 
Innovative systems and methods for visualizing information collected from 
analyzing samples are provided. The samples may include nucleic acids, proteins, or 
other polymers. Gene expression level as determined from analysis of a nucleic acid 
sample is one possible analysis result that may be visualized. In one embodiment, a 
computer system may display the expression levels of mxxltiple genes simultaneously in a 
way that facilitates user identification of genes whose expression is significant to a 
characteristic such as disease or resistance to disease. Additionally, the computer system 
may facilitate display of further information about relevant genes once they are identified. 
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DEFINITION 
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NID 

KEYWORDS 
SOURCE 

ORGANISM 



REFERENCE 
' AUTHORS 

TITLE 

JOURNAL 

COMMENT 



Isobe^M., Kumura,Y., Ogita^Z., Hinoda^Y., 



of a human 



1607-1615 



Phone : 
Fax: 



FEATURES 
x: source 



gene 



CDS 



p o 1 yA__s i gn al 
polyA_site 

BASE COUNT 652 

ORIGIN 



HUMLCPTP 2691 bp mRNA PRI 06-NOV-1992 

Hioinan mRNA for protein-tyrosine phosphatase, complete cds. 
D11327 
g219901 

protein-tyrosine phosphatase. 

Human T cell, lambda-gtlO library, cDNA to mRNA, 
Homo sapiens 

Eukaryotae; mitochondrial eukaryotes; Metazoa; Chordata; 
Vertebrata; Mammalia; Eutheria; Primates; Catarrhini; Hominidae; 
Homo . 

1 (bases 1 to 2691) 
Adachi,T, , Sekiya,M, 
Imai,K. and Yachi,A. 

Molecular cloning and chromosomal mapping 
protein-tyrosine phosphatase LC-PTP 

Biochemical and Biophysical Research Communication 186, 
(1992) 

Submitted (22-MAY-1992) to DDBJ by: 

Masaaki Adachi 

Sapporo Medical College 

S1W16 Chuo-ku 

Sapporo 060 

Japan 

011-611-2111 
011-613-1141. 
Location/Qualifiers 
1. .2691 

/organism="Homo sapiens" 
/db_xref="taxon: 9606" 
/cell_type="T cell" 
/clone_lib="lambda-gtlO" 
105.. 1187 
/gene="LC-PTP" 
105. .1187 
/gene="LC-PTP" 
/ codon_s tart=l 

/product="protein-tyrosine phosphatase" 
/db_xref ="PID : dl002425" 
/db_xref="PID:g219902" 

/ trans lation= "MVQAHGGRSRAQPLTLSLGAAMTQPP PEKTPAKKHVRLQERRGS 
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KIPSNFVSPEDLDIPGHASKDRYKTILPNPQSRVCLGRAQSQEDGDYINANYIRGYDG 
KEKVYIATQGPMPNTVSDFWEMVWQEEVSLIVMLTQLREGKEKCVHYWPTEEETYGPF 
QIRIQDMKECPEYTVRQLTIQYQEERRSVKHILFSAWPDHQTPESAGPLLRLVAEVEE 
SPETAAHPGPIWHCSAGIGRTGCFIATRIGCQQLKARGEVDILGIVCQLRLDRGGMI 
QTDEQYQFLHHTLALYAGQLPEEP S P " 
444.. 1172 
/gene="LC-PTP" 

/note=" single catalytic domain" 
2667.. 2672 
2691 

a 816 c 707 g 516 t 



misc feature 



1 ggagacagac agacagctgg caagaggcag cctgggggcc acagctgctt cagcagacct 
61 catggctgag tgagcctccc ctgggcccag caccccacct cagcatggtc caagcccatg 
121 gggggcgctc cagagcacag ccgttgacct tgtctttggg ggcagccatg acccagcctc 
181 cgcctgaaaa aacgccagcc aagaagcatg tgcgactgca ggagaggcgg ggctccaatg 
241 tggctctgat gctggacgtt cggtccctgg gggccgtaga acccatctgc tctgtgaaca 
301 caccccggga ggtcacccta cactttctgc gcactgctgg acaccccctt acccgctggg 
361 cccttcagcg ccagccaccc agccccaagc aactggaaga agaattcttg aagatccctt 
421 caaactttgt cagccccgaa gacctggaca tccctggcca cgcctccaag gaccgataca 
481 agaccatctt gccaaatccc cagagccgtg tctgtctagg ccgggcacag agccaggagg « 
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